Papers by Špela Arhar Holdt

4 papers
Gigafida 2.0: The Reference Corpus of Written Standard Slovene (2020.lrec-1)

Copied to clipboard

Challenge: Gigafida reference corpus of Slovene is updated with new material and tools . focus of upgrade was on transformation from general reference corp to standard reference corp .
Approach: We present a new version of the Gigafida reference corpus of Slovene . the upgrade includes new material and better tools for annotating it .
Outcome: The new version of the Gigafida reference corpus of Slovene is described . the whole Gigido corpus was deduplicated for the first time .
Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning (2020.lrec-1)

Copied to clipboard

Challenge: Lack of wide-coverage and high-quality LRs is a longstanding issue in natural language processing (NLP) however, there are no large initiatives of similar scale for creating new LR or improving existing ones.
Approach: They propose a generic approach to combine implicit crowdsourcing and language learning to mass-produce language resources (LRs) they describe its core paradigm that consists in pairing specific types of LRs with specific exercises .
Outcome: The proposed approach can be used in several learning scenarios to produce a multitude of NLP resources and alleviate the long-standing issue of the lack of LRs.
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)

Copied to clipboard

Challenge: a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years.
Approach: They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene.
Outcome: The revised corpus, built on its predecessor, doubles in size and depth of annotation layers.
Towards an Ideal Tool for Learner Error Annotation (2024.lrec-main)

Copied to clipboard

Challenge: 'correction annotation' is a technique that has been used for many years to correct errors in learner corpora.
Approach: They propose to use SVALA to annotate and analyse corrections in learner corpora using a parallel aligned approach to visualisation and annotation.
Outcome: The proposed tool supports multiple annotation systems, localisation into other languages, and the development of more complex annotation systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations